Back

The Annals of Applied Statistics

Institute of Mathematical Statistics

Preprints posted in the last 30 days, ranked by how well they match The Annals of Applied Statistics's content profile, based on 19 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Recalibrating Mendelian randomization under winner's curse, sample structure and polygenicity

Yang, Y.; Lin, Z.; Xue, H.; Zhu, X.

2026-07-07 genetic and genomic medicine 10.64898/2026.06.25.26356593 medRxiv
Top 0.1%
6.6%
Show abstract

Recently, Hu et al. (2024) conducted a benchmarking study showing that most existing Mendelian randomization (MR) methods exhibit substantial bias and inflated type-I error rates in real data. They attributed these failures to two largely neglected sources of bias: winner's curse and polygenicity-induced bias. Although a few methods have been developed to address one or both of these issues, existing approaches either do not fully account for both biases or are restricted to the univariable setting. In this paper, we propose a multivariable Rao-Blackwellization that corrects winner's curse while accounting for polygenicity and sample structure in a unified framework. Unlike univariable Rao-Blackwellization, where instrument selection yields a truncated normal statistic amenable to a Mills-ratio correction, multivariable Rao-Blackwellization conditions on a noncentral $\chi^2$ statistic, for which no analogous correction is available. We derive closed-form conditional moments under this instrument selection model and use them to construct bias-corrected summary statistics that can be integrated into a wide range of existing MR methods. Simulations and real data analyses show that, when combined with methods such as MR-cML and MR-BEE, the proposed correction substantially improves type-I error control and yields more robust inference.

2
Model-based inference of gene expression noise from single-cell RNA-sequencing data

Giersdorf, F.; Rogers, D. W.; Christensen, S.; Dutheil, J. Y.

2026-06-23 bioinformatics 10.64898/2026.06.18.733122 medRxiv
Top 0.1%
4.4%
Show abstract

The heterogeneity of expression levels among genetically identical cells, termed gene expression noise, is a property of the gene expression process whose importance in the biology of organisms and their evolution is increasingly recognized. Measuring gene expression noise requires single-cell expression data, as obtained from single-cell RNA sequencing (scRNASeq). Its estimation, however, is challenging owing to (i) the presence of technical noise in addition to biological noise, and (ii) the heterogeneity of cell types in the sampled population. We propose a maximum-likelihood framework to infer biological noise from scRNASeq data, while accounting for technical noise, dropout probabilities, and distinct cell sequencing depths. We demonstrate the parameter identifiability using simulations and that the resulting noise estimates are uncorrelated from the mean gene expression, and therefore do not need extra correction in downstream analyses, easing intra- and inter-genome comparisons. Using two technical replicates of scR-NASeq data from the wild yeast Saccharomyces paradoxus, we show that expression noise can be inferred in a reproducible manner.

3
Calibrating machine learning approaches for probability estimation without calibration data

Di Carluccio, E.; Koliopanos, G.; Ojeda, F. M.; Weimar, C.; Ziegler, A.

2026-07-13 epidemiology 10.64898/2026.07.10.26357723 medRxiv
Top 0.1%
3.1%
Show abstract

Statistical prediction models for binary outcomes are becoming increasingly popular. One significant challenge is calibrating these models to suit the characteristics of a target population that is structurally different from the original population. Calibration is especially challenging when there is no training data available from the target population. To address this problem, we propose a novel calibration method, SimCal, which uses synthetic data generated from the model development data in conjunction with marginal statistics from the calibration cohort. We show that expert judgment modeling (EJM) may be used for calibration if cross-sectional data from the target population are available comprising expert judgments about the potential outcome and the covariates. We describe three alternative calibration approaches when calibration data are lacking: similarity-binning averaging (SBA), adaptive calibration of predictions (ACP), and Elkan calibration. In a simulation study, we compare SBA, ACP, Elkan calibration, and SimCal. R code for applying these methods is provided from the re-analysis of data on coronary artery disease. We illustrate all 5 calibration approaches with a real data set for predicting functional outcome after stroke and all approaches but EJM in the re-analysis of the Cleveland Clinic data. None of the approaches performed convincingly well in all situations. SimCal performed well when model parameters were correctly specified. EJM failed on the stroke data. Further research is urgently required for calibration in the absence of calibration data.

4
Overinflation and overconcentration: why Cauchy perturbation kernels are the right choice for ABC-SMC

Sturrock, M.; Shahrezaei, V.

2026-07-09 systems biology 10.64898/2026.06.24.734205 medRxiv
Top 0.1%
2.1%
Show abstract

Approximate Bayesian computation sequential Monte Carlo (ABC-SMC) propagates its particles with a perturbation kernel, and with the standard Normal kernel it degrades sharply as the parameter dimension grows, a failure usually attributed to dimension itself. We show instead that it is governed by the quality of the summary statistics, with dimension entering only through a separate and milder mechanism, and that the two must act together for the Normal kernel to break. The first ingredient is covariance overinflation: the kernel covariance, estimated from the particle cloud, overshoots the true posterior covariance by a factor set by information loss in the summary statistics. We derive this overscaling factor in closed form for a Gaussian model with sufficient statistics and show that it stays modest at any dimension, shrinking toward its baseline value as the tolerance tightens; the extreme values seen in practice (of order 103) are a signature of insufficient summaries, not of dimension. The second ingredient is perturbation overconcentration: the normalised Normal step size concentrates around one as the dimension grows, so every proposal overshoots by the same factor. Either ingredient alone is harmless; only their combination breaks the Normal kernel. A Cauchy kernel (multivariate t with one degree of freedom) removes the concentration, keeping a positive acceptance rate under arbitrary overscaling at a bounded worst-case cost of 1.87x in expected squared jump distance. In a Metropolis-Hastings framework we derive closed-form acceptance rates for both kernels that illustrate the advantage of the Cauchy kernel in this limit. A series of full ABC-SMC computational experiments on five problems at d = 12, including a hierarchical gene-expression model, show the Cauchy reducing the sliced Wasserstein distance to the reference posterior by factors of up to 50 with the same simulation budget. Since the summary statistics are commonly insufficient for the models that require ABC, overinflation is structural and the Cauchy perturbation kernel is the right default for problems in higher dimensions.

5
Scalable multi-group nonnegative spatial factorization for spatial genomics data with cell-type heterogeneity

Chumpitaz-Diaz, L.; Shrestha, P.; Engelhardt, B. E.

2026-07-03 genomics 10.64898/2026.06.29.735224 medRxiv
Top 0.1%
1.7%
Show abstract

Spatial transcriptomics (ST) technologies enable the study of gene expression within the spatial context of tissues, providing insights into tissue structure, cellular interactions, and disease progression. However, existing dimension reduction methods often overlook spatial information or struggle to distinguish spatial gene patterns from those driven by cell-type differences, limiting biological interpretability by convolving differences in gene expression patterns with differences in cell-type proportions. To address these challenges, we introduce the scalable multi-group nonnegative spatial factorization (smNSF), a computationally-tractable probabilistic framework that integrates spatial coordinates and cell-type labels into a unified matrix factorization model. By using multi-group Gaussian processes (MGGPs) as priors, our model captures complex spatial variation in a cell-type specific way while enforcing nonnegativity to enhance interpretability. We develop a variational inference framework for MGGPs that supports scalable optimization and improves the numerical stability of smNSF. Across seven spatial transcriptomics datasets spanning diverse technologies and tissues, smNSF recovers sparse, interpretable spatial factors and, through its cell-type conditional posteriors, organizes them into cell-type enriched, cell-type specific, and universal spatial programs that are not apparent from marginal factors alone. Given cell-type labels in ST data, smNSF enables cell-type aware spatial decompositions and supports cell-type conditional posteriors for in silico exploration of relationships between spatial patterns and cellular identity.

6
GCBM-DCT-HV-Bio-NL-Grow-CHG-CSM-RHEC: A Unified Geometric, Biological, Causal, and Regenerative Framework for Mechanism-Aware Tissue and Connectome Modeling

Xu, T.; Hu, Z.; Sun, X.; Jin, L.; Xiong, M.

2026-06-29 bioinformatics 10.64898/2026.06.24.734320 medRxiv
Top 0.1%
1.5%
Show abstract

Modern biological prediction problems increasingly require models that go beyond Euclidean feature regression and local graph smoothing. Tissue, cellular, and connectome systems are nonlinear, geometry-dependent, intervention-sensitive, history-dependent, and subject to regenerative or homeostatic constraints. We propose GCBM/DCT/HV/Bio/NL/Grow/CHG/CSM/RHEC, a unified model for mechanism-aware biological prediction. The model integrates geometric connectome dynamics, differentiable charted tissue geometry, Hamiltonian latent transport, nonlinear biological kinetics, nested latent memory, continual growth without overwriting, causal hypergraph structure, causal structure modeling, and regenerative homeostatic error correction. Unlike Euclidean baselines, which treat observations as flat vectors, and local graph baselines, which use neighborhood smoothing without mechanistic structure, the proposed model represents biological states (Trapnell 2015) as coupled geometric, dynamical, causal, and regenerative objects. We evaluate the model on four synthetic toy studies, Toy A, B,C, D, designed to reflect increasing biological complexity: local Euclidean structure, nonlinear mechano-chemical interaction, causal intervention response, and out-of-distribution regenerative shift. Compared with Euclidean and local graph baselines, the full model achieves the lowest mean squared error across all four toy studies. Relative to the Euclidean baseline, the full model reduces MSE by approximately 63.0%, 89.1%, 89.0%, and 90.9% on Toy A, Toy B, Toy C, and Toy D, respectively. These results support the value of integrating geometry, mechanism, causal structure, adaptive growth, and regenerative correction into a single predictive architecture (Figure 1).

7
BOSE: A Bayesian Order Statistics-Based Estimator for Recovering the Sample Mean and Standard Deviation

Pan, W.; Lu, Z.; Jiang, W.; Lim, J.; Xu, L.; Wang, X.

2026-07-01 bioinformatics 10.64898/2026.06.26.734829 medRxiv
Top 0.1%
1.5%
Show abstract

In meta-analyses of continuous outcomes, the sample mean and standard deviation (SD) are essential for synthesizing effect sizes across studies. However, clinical studies frequently report alternative summary statistics, such as the median, quartiles, and range. To enable inclusion of such studies, various methods have been proposed to estimate the sample mean and SD from these reported summaries. We propose the Bayesian Order Statistics-based Estimator (BOSE), which leverages the joint likelihood of observed order statistics together with weakly informative priors to obtain the full posterior distribution for the mean and SD without relying on computationally intensive iterative procedures such as Markov chain Monte Carlo algorithms. Our numerical studies demonstrate that BOSE performs competitively with existing approaches in estimating the mean, while achieving superior performance for estimating the SD across all evaluated scenarios, particularly in small-sample settings. Under non-normal distributions including skewed, heavy-tailed, and bimodal settings with mild or moderate deviations from normality, BOSE remains robust and stable, whereas methods specifically designed for skewed distributions may become unstable or even inapplicable. Beyond point estimation, BOSE naturally provides empirically validated posterior credible intervals, enabling researchers to formally quantify uncertainty for study-level estimates and make reliable, evidence-based decisions in meta-analytic research synthesis. A publicly accessible web application implementing BOSE and competing methods is also provided to facilitate practical use in meta-analytic research.

8
Lineage-aware stochastic modeling reveals gene-expression dynamics in development and disease

Xing, J.; Staklinski, S. J.; Liu, Z.; Nowak, D.; Siepel, A.

2026-06-28 bioinformatics 10.64898/2026.06.25.734628 medRxiv
Top 0.1%
1.4%
Show abstract

Gene expression evolves dynamically along cell lineages, yet most analysis methods treat single-cell RNA-seq (scRNA-seq) data as static snapshots and fail to exploit phylogenetic relationships among cells. Recent advances in cell-lineage tracing now enable the reconstruction of high-resolution lineage phylogenies, providing a natural framework for identifying when and where transcriptional changes arise during development, differentiation, and disease progression. Some models of gene expression have begun to consider phylogenetic structure, but they generally rely on imprecise Gaussian assumptions, focus on endpoint-level comparisons, or fail to consider sparse and overdispersed scRNA-seq read counts. Here, we present LaVOUS (Lineage-aware Variational Ornstein-Uhlenbeck Single-cell RNA-seq analysis), a probabilistic framework that couples lineage-based models of latent dynamics derived from the Brownian motion and Ornstein-Uhlenbeck stochastic processes with negative-binomial observation models and scalable variational inference. LaVOUS enables likelihood-based tests for cellular heritability and branch-specific shifts in gene expression, as well as phylogenetic reconstruction of latent expression histories. In simulations, LaVOUS outperformed Gaussian method in detecting lineage-associated expression changes and produced accurate reconstructions of expression histories across expression levels. We additionally applied LaVOUS to paired single-cell lineage and transcriptomic data from metastatic lung cancer, class-switching B cells, and the developing brain. Across these settings, LaVOUS identified lineage-associated expression changes related to metastatic progression, B-cell isotype switching, and dopaminergic and glutamatergic neuron differentiation. By providing an expressive framework for modeling sparse count data on lineage trees, LaVOUS establishes a foundation for studying single-cell expression dynamics across developmental and disease contexts, with natural extensions to multi-gene regulation, lineage uncertainty, and multi-modal integration.

9
netPCF: Geometry-Aware Pair Correlation Functions for Spatial Biology

Moore, J. W.; Bull, J. A.; Byrne, H. M.

2026-07-07 bioinformatics 10.64898/2026.07.02.736020 medRxiv
Top 0.2%
1.1%
Show abstract

Spatial organisation is a defining feature of biological systems, underpinning cellular interactions, tissue function, disease progression and therapeutic response. Identifying and quantifying spatial organisation may require methods that resolve relationships across spatial scales. The pair correlation function (PCF) quantifies spatial dependence between points across multiple length scales, but its standard Euclidean formulation is poorly suited to data defined on irregular, curved or otherwise structured domains, where tissue geometry may constrain biological organisation and distort Euclidean distances. Here, we introduce netPCF, a geometry-aware extension of the PCF for quantifying spatial organisation on complex biological domains. By representing tissue structures, anatomical surfaces and other constrained geometries as spatial networks, netPCF generalises the PCF beyond extrinsic Euclidean settings. The framework derives the expected behaviour of the statistic under complete spatial randomness using interpretable finite-support kernels, provides bootstrap-based uncertainty quantification, and includes practical criteria for assessing domain discretisation adequacy. We further extend netPCF to marked (labelled) biological data using feature kernels for categorical and continuous attributes, enabling unified analysis of cell identities, marker intensities, phenotypic states, gene expression and other quantitative features on structured domains in any spatial dimension. All methods are implemented in the open-source Python package spacenet. Synthetic studies show that netPCF recovers classical Euclidean behaviour on sufficiently resolved networks and is robust to common imaging noise. We demonstrate its utility in two biological applications. In three-dimensional imaging mass cytometry data from HER2+ breast carcinoma, netPCF separates tissue architecture-driven proximity from biologically meaningful endothelial and immune cell organisation. In reconstructed surfaces of developing murine embryos, netPCF identifies a transition in the Wnt1-Wnt6 relationship from short-range co-localisation at E9.5 to spatial exclusion at E11.5, a pattern of ectodermal boundary refinement not captured by prior voxel-wise co-expression analysis. Overall, netPCF provides a statistically grounded and practical framework for quantifying spatial organisation on complex biological domains.

10
Two-Sample Instrumental Variables under Population Mismatch: A Transportability Framework with Bias Diagnostics

Qian, Y.; Song, Y.

2026-06-26 epidemiology 10.64898/2026.06.15.26355602 medRxiv
Top 0.2%
1.0%
Show abstract

Instrumental variable (IV) methods are widely used in health and social sciences to estimate causal treatment effects among compliers. In certain research settings, the instrument-treatment association (first stage) and the instrument-outcome association (reduced form) are each estimated from a different dataset. Two-Sample Instrumental Variables (TSIV), proposed by Angrist and Krueger (1992), addresses this by combining first-stage and reduced-form estimates from separate data sources into a single causal effect estimate. However, TSIV identification requires that instrument compliance behavior be consistent across the two samples, a condition that is rarely verified in practice. We show mathematically and empirically that when compliance differs between samples, the raw TSIV estimator does not converge to the true Local Average Treatment Effect (LATE) and instead attenuates toward a predictably biased limit proportional to the ratio of first-stage compliance rates between the two samples. To address this, we formalize a framework for estimating LATE with TSIV under two key assumptions: (1) Covariate Overlap, requiring that the two samples share sufficient common support in their covariate distributions, and (2) Compliance Transportability, requiring that compliance behavior is identical across populations after conditioning on observed covariates. We consider a setting in which a health policy instrument and outcomes are recorded in administrative claims while treatment and covariates are collected in a survey. We use a C-statistic derived from pooled covariates to detect population mismatch and an Inverse Probability Weighting (IPW) correction that reweights the first-stage sample to approximate the administrative covariate distribution. In Monte Carlo simulations across eight scenarios calibrated to a survey-Medicaid setting, IPW-TSIV reduces bias in estimating the LATE, achieving 88% reduction in the primary scenario, 82% under severe selection, and 79% when state-level expansion policy drives compliance heterogeneity. We further validate this framework using the Oregon Health Insurance Experiment, where partitioning the public-use lottery data (N = 24,646) into two non-overlapping samples with substantively meaningful compliance heterogeneity yields a verifiable benchmark against the true causal effect. IPW-TSIV reduces mean absolute bias by 71.6% relative to the oracle S2-specific LATE across 10 independent replications (C-statistic = 0.78), outperforms naive TSIV in all 10 splits, and reduces mean bias relative to the full-data LATE from +0.016 to +0.008. This framework provides applied researchers with actionable diagnostic thresholds to detect sample mismatch, validate transportability assumptions, and determine when structural TSIV estimation is reliable.

11
When Less Is Not More: DICEPro Mitigates the Impact of Incomplete Reference Matrices on Cellular Frequency Deconvolution.

BA, K.; Thiebaut, R.; Hinaut, X.; Hejblum, B. P.

2026-06-22 bioinformatics 10.64898/2026.06.17.732876 medRxiv
Top 0.2%
0.9%
Show abstract

Cellular deconvolution aims to estimate the frequencies of different cell populations from gene expression measurements in a biological sample. Supervised approaches, such as CIBERSORTx and DISSECT, critically depend on the reference signature matrix, which encodes the gene expression profiles of cell-types based on prior knowledge. Despite numerous deconvolution methods, the impact of missing cell populations in the reference matrix remains understudied. Here, we evaluate the robustness of state-of-the-art deconvolution approaches using simulations based on real dataset examples combined with statistical modeling, validated against published data, and multiple real benchmark datasets. Results show that deconvolution performance remains stable when the reference matrix includes most cell-types, but declines sharply as the matrix becomes incomplete, especially for abundant cell populations. To address the limitations of incomplete reference matrices, we introduce DICEPro, an optimization-based framework designed to enhance existing deconvolution methods. By systematically adjusting the reference signatures, DICEPro better accounts for missing or underrepresented cell populations, leading to improved precision and robustness. We show that DICEPro consistently boosts deconvolution performance across both simulated datasets, derived from real data examples, and multiple real biological datasets, offering a practical solution when standard methods are hindered by incomplete references.

12
Neural Processes with Normalizing Flows for Wheat Height Estimation

Boss, M.;Volpi, M.;Roth, L.

2026-07-09 Plant Biology 10.64898/2026.06.24.734247 medRxiv
Top 0.2%
0.9%
Show abstract

In this work, we investigate modeling plant traits over time using neural processes, a class of machine learning models that learn distributions over functions. Plant growth is an inherently stochastic process with complex dynamics measured mostly at irregular times throughout the growing seasons. While individual trait trajectories may be simple, their distributions are shaped by complex interactions between genotype, environment, and other factors. In particular, we focus on plant height in wheat, a deceptively simple-looking trait with complex dynamics. To model these trajectory distributions, we evaluate neural processes and in particular extensions using normalizing flows, with different combinations of genotype and environmental covariates. For controlled evaluations, we generate synthetic wheat height trajectories calibrated against Swiss weather station records and the FIP1 dataset. To fully evaluate these trajectory distributions, we use signatures, vector representations of sequential data, together with Sig-MMD and the recently introduced CSig-MMD. Sig-MMD enables direct pathwise comparison of predicted and simulator trajectory distributions, while CSig-MMD focuses this comparison on the tail, including lodged trajectories. Together, these metrics allow us to assess whether the models capture the full distribution of growth trajectories, including rare outcomes.

13
Fast Diffusion of Bound Ca: Analytical and Experimental Characterization of One- and Two-Dimensional Traveling Waves

Mironov, S.

2026-07-10 biophysics 10.64898/2026.07.06.735233 medRxiv
Top 0.2%
0.9%
Show abstract

Reaction diffusion (RD) systems play a fundamental role in numerous biochemical and biophysical processes. Here, we present a novel analytical framework for solving RD equations by applying the Wentzel Kramers Brillouin Jeffreys (WKBJ) formalism to Ca nanodomains generated by individual membrane channels, a widely used paradigm for intracellular Ca signaling. Previous models have primarily focused on stationary Ca nanodomains while neglecting diffusion and saturation of intracellular Ca buffers and sensors. In contrast, we derive analytical solutions without these simplifying assumptions. Our analysis demonstrates that sustained Ca influx generates continuously expanding distributions of free Ca, whereas Ca bound buffers and sensors propagate as traveling waves. These predictions are supported experimentally by measurements of one-dimensional fluorescence profiles produced by single-channel activity and two-dimensional profiles generated by whole cell Ca currents. The analytical framework developed here readily extends Michaelis Menten type kinetics to reaction diffusion systems and may therefore be broadly applicable to biochemical and biophysical processes in which diffusion cannot be neglected.

14
Uncovering internal states with a robust shared-state multi-neuron GLM-HMM framework

Lawrence, A.; Yezerets, E.; Janak, P. H.; Charles, A.

2026-07-02 neuroscience 10.64898/2026.06.27.734988 medRxiv
Top 0.2%
0.8%
Show abstract

Neural systems exhibit multiple firing states that reflect an organism's internal state and modulate the relationship between external environmental stimuli and behavior. Several studies have inferred these latent states by supplementing the traditional hidden Markov Model (HMM) with generalized linear models (GLMs) with non-Poisson behavioral observations. However, understanding the relationship between internal brain states and behavior also requires modeling the neural activity. Nonetheless, fitting multi-neuron GLM-HMMs is non-trivial due to high sparsity, collinearity, and low trial counts in neuronal datasets. Therefore, we built a robust multi-neuron GLM-HMM framework that uncovers latent states from population activity while incorporating the influence of time-stamped task variables and spike histories. To obtain reliable model parameters, we employ a modified expectation-maximization procedure. Specifically, we show that incorporating neuron-adaptive penalization in the maximization step overcomes the covariate co-linearity issues typical of time-stamped events and sparse spiking, yielding stable estimates of Poisson GLM coefficients. Furthermore, we incorporate a trust-region algorithm to ensure stable M-step convergence in the presence of ill-conditioned Hessians that can lead to unstable Newton-Raphson updates. We further demonstrate the utility of leave-one-out cross-validation analysis for evaluating model performance on datasets with low trial counts and without breaking their temporal structure. We evaluate our framework on three electrophysiological datasets from primates and rodents as they perform a decision-making task, demonstrate stable model convergence, and discuss the behavioral relevance of the inferred states.

15
Statistical Inference and Power Analysis for Comparative F1 and Fβ Scores under Correlated Classifier Pairs

Hsu, C.-Y.; Liu, Q.; Shyr, Y.

2026-07-17 dermatology 10.64898/2026.07.15.26358166 medRxiv
Top 0.2%
0.8%
Show abstract

As machine learning and artificial intelligence systems are increasingly used in healthcare, rigorous evaluation of their classification performance has become critical. The F1 and F{beta} scores are widely adopted metrics for assessing performance in imbalanced biomedical data. Recently, we introduced psF1, a unified statistical framework for inference and study design for single and comparative F1 and F{beta} scores under the assumption of independent classifiers. In practice, however, benchmarking two classifiers on the same dataset creates a correlated paired setting. Ignoring this intrinsic dependency leads to overestimation of the standard error and a substantial loss of statistical power. To address this, we develop psF1pair, an advanced framework for statistical inference and power analysis that explicitly accounts for correlations between classifier pairs. Extensive simulation studies demonstrate the performance of psF1pair, and its utility is further illustrated through application to a real-world imaging classification system. As expected, higher correlation between classifiers yields narrower confidence intervals and enhanced statistical power. A freely available R package is provided to facilitate implementation, supporting accurate evaluation and study design for predictive and classification models in biomedical research.

16
Why linkage disequilibrium measures disagree: Fisher geometry of rare common haplotype structure

Ichikawa, Y.

2026-07-07 genetics 10.64898/2026.07.02.736022 medRxiv
Top 0.2%
0.8%
Show abstract

Conventional LD measures such as r2 perform poorly in the rare common regime, particularly in asymmetric configurations such as nested haplotype structure. Because r2 is symmetric and quadratic, it removes directional structure in two ways: squaring discards the sign, or phase, retained by the signed LD coefficient D, while symmetric normalization hides the asymmetry between the conditional probabilities P(A|B) and P(B|A). Although D recovers the phase, it is locus symmetric and unnormalized; its magnitude is hard to compare across frequency regimes and it does not by itself express which way the asymmetry runs. We therefore analyze the conditional-probability asymmetry {Delta} = P(A|B) - P(B|A), together with r2 and D, as distinct scalar functions on the haplotype simplex under the Fisher information metric. The conditional probabilities P(A|B) and P(B|A) are bounded in [0, 1], directly express carrier-set inclusion, and are more readily visualized than D. Moreover, their difference admits the exact decomposition {Delta} = M + C into a marginal frequency term M and an LD-coupled term C. Prior work has characterized either the mathematical behavior of LD normalizations across allele-frequency space or the Fisher geometry of the haplotype simplex, but not their connection. We bridge this gap by showing that the geometric structure of the simplex explains why LD measures disagree in the rare common regime and why symmetric normalizations such as r2 lose directional information. We show that the fixed-frequency leaf is intrinsically anisotropic, positively curved, and frequency-dependent under the Fisher metric. These geometric predictions are tested empirically , in phased 1000 Genomes data1 and a two locus Wright Fisher model, in a companion paper (Ichikawa, preprint); the present note develops the geometry itself. Keywords: linkage disequilibrium; Fisher information metric; haplotype simplex; rare variant; conditional-probability asymmetry; nested haplotype structure

17
Model-optimized stimulus distortions for adaptive estimation of individual sensory representations

Casco-Rodriguez, J.; Hong, F.; Brainard, D. H.; Feather, J.; Lipshutz, D.

2026-07-08 neuroscience 10.64898/2026.07.02.736141 medRxiv
Top 0.2%
0.6%
Show abstract

Representations of the same physical stimulus vary between individuals. Characterizing individual differences has practical implications, but is challenging because these representations are not directly observable. Given a model of how representations vary within a population, we propose a Bayesian adaptive procedure for estimating an individual observer's representation from a series of targeted perceptual discrimination judgments. A key component of our approach is using Fisher information to identify stimulus distortions that efficiently differentiate observers in the population. As a proof of concept, we focus on individual differences in color perception and simulate observers with cone fundamentals drawn from an individual colorimetric observer model. We demonstrate that our approach can recover key aspects of a sampled observer's cone fundamentals using simulated three-alternative forced-choice oddity judgments with approximately 500 trials, corresponding to an experimental duration of approximately one hour. Our Bayesian adaptive framework provides a promising and generalizable approach to efficiently link behavioral measurements to individual differences in sensory representations.

18
G-LATO: Inference of Spatial Latent Ordering via Deep Gaussian Processes

Zago, M.; Mukherjee, S.; Schleicher, J. T.; Bürkner, P.; Tabatabai, G.; Claassen, M.

2026-06-29 bioinformatics 10.64898/2026.06.23.734031 medRxiv
Top 0.2%
0.6%
Show abstract

Spatial transcriptomics enables the study of cells within their native tissue context, yet identifying gradients of cellular development remains challenging. We introduce a deep Gaussian process model to address this gap. Our method recovers spatially smooth gradients explaining observed gene expression. We illustrate our method on healthy liver and glioblastoma data in reconstructing known spatial organisation and uncovering new pathological gradients, thus providing robust inference for spatial biology.

19
Linking Systemic Endotoxin Exposure to Retinal Microglia Migration Through Mathematical Modeling

Jackson, T. M.; Cassidy, T.; Dando, S. J.; Jenner, A. L.

2026-07-05 immunology 10.64898/2026.06.30.735687 medRxiv
Top 0.2%
0.6%
Show abstract

Microglia are the resident immune cells of the central nervous system (CNS), including the brain, spinal cord, and retina, where they serve as the first line of defense against infection and inflammation. Dysregulated microglia activity has been implicated in vision-threatening diseases, highlighting the need to understand how retinal microglia respond to inflammatory stimuli. Importantly, acute inflammation induces substantial redistribution of microglia across retinal layers, yet the mechanisms governing this migration remain poorly understood. Here, we develop the first mathematical model of retinal microglia migration during inflammation to determine how inflammatory exposure, administration route, and species-specific pharmacokinetics shape redistribution dynamics across the retina. The model couples lipopolysaccharide (LPS) pharmacokinetics with microglia migration between the outer plexiform layer (OPL), inner plexiform layer (IPL), and ganglion cell layer/nerve fiber layer (GCL/NFL). Model parameters are calibrated to retinal microglia density measurements from mice following LPS (bacterial endotoxin) challenge, before extending the framework to rats and rhesus macaques to investigate species-specific responses. Simulations also compare how administration route, i.e. intravenous or intraperitoneal injections, alter retinal LPS exposure and subsequent microglia redistribution. Our results suggest that redistribution patterns are driven primarily by LPS delivery route and species-specific pharmacokinetics, rather than the initial microglia distribution across retinal layers. Together, these findings provide new insight into immune cell reorganization in the inflamed retina and demonstrate how mechanistic mathematical modeling can be adapted across experimental designs, administration routes, and animal species.

20
Proliferative and Motile Cell Interplay in Glioma Invasion: Go-or-Grow Switching Caps the Invasion Speed

Sadhukhan, S.; Santra, D.

2026-07-07 biophysics 10.64898/2026.07.01.735477 medRxiv
Top 0.2%
0.6%
Show abstract

Diffuse gliomas are deadly because the individual tumor cells invade - they travel far from the imageable mass, so it is impossible to remove the tumor completely. On the cellular level, glioma cells seem to be in either a "go" state (in which they do not divide) or a "grow" state (in which they do not migrate). We investigate what this tiny choice has to say about the large-scale speed of the invasion front and whether the implication is sufficiently strong to rule out the classical description of the Fisher-Kolmogorov-Petrovsky-Piskunov (Fisher-KPP) type, in which a single phenotype migrates and proliferates. We derive a two-phenotype reaction-diffusion model with density-dependent switching, and we prove the cooperative (quasi-monotone) structure and the associated comparison principle and study travelling-wave solutions of the model. A leading-edge linearization gives minimal front speed as minimizer of an explicit dispersion relation, and direct simulation verifies the predicted speed. In the experimentally relevant fast switching limit, we find a closed-form expression for the speed, that is, we obtain an effective Fisher-KPP equation with rescaled diffusivity and growth rate, with the fractions of the phenotypes. The "go-or-grow" (GoG) front can move at a maximum speed of half the Fisher speed for the same single-cell motility $D$ and proliferation rate $r$, which occurs only when the cells divide their time equally between the two phenotypes. This bound is directly testable: measurement of the front speed, plus independent determination of $D$ and $r$, discriminates the two hypotheses, and in the GoG case, yields recovery of the phenotype balance. We then extend the result to anisotropic (DTI-informed) invasion along white-matter tracts and discuss implications for understanding clinical measurements of growth rate.